Zhipu releases open-source GLM-5.3-Flash multimodal model
The new 320B-parameter model activates 18B parameters and is positioned as a cheaper, faster member of the GLM-5 family.

Chinese AI company Zhipu has released GLM-5.3-Flash, described as the first native multimodal model in its GLM-5 family.
The model has 320 billion total parameters but activates 18 billion at a time, a mixture-of-experts approach designed to provide large-model capability without running the entire network for every request.
Native multimodal input
GLM-5.3-Flash is intended to handle more than text, bringing image understanding into the model’s core capabilities. Zhipu is releasing it as an open-source model for developers to test and integrate.
IT Home reports that the model is being offered at a limited-time price around one twentieth of GLM-5.3, making cost a major part of the launch pitch.
Why the parameter count is complicated
The headline 320B number sounds enormous, but the 18B active-parameter figure is closer to the amount of computation used for an individual response. That does not make the model small, but it does make the architecture more efficient than a dense 320B system.
Real-world speed, hardware requirements and benchmark quality will matter more than the parameter figure alone.
Our opinion
The interesting part of GLM-5.3-Flash is not simply that it is big. It is that open models are increasingly competing on efficiency, modality and price rather than treating scale as the only scoreboard.
That discount could attract developers who want multimodal capability without committing to the cost of a flagship model. The usual warning applies: cheap access is only useful if the output is reliable enough to use.
If Zhipu can pair the architecture with genuinely strong image and text performance, GLM-5.3-Flash could be a practical workhorse rather than another giant model people admire from a safe distance.